Tencent Hunyu2026-09-15 09:11:00EvolveScaler tests whether LLMs can keep up with changing information, with the best GPT-5.5 score at 59.3%Teams including Tencent Hunyuan have released EvolveScaler, a benchmark built to test whether large language models can reason over information that keeps changing over time. The setup injects updates, reversals, additions, and invalidations into long passages, then asks models to answer based on the latest state rather than retrieve a ready-made sentence from the text. One example starts with 40 days of game records and then asks a counterfactual question about what would happen if a small boss had not been defeated on day 7, forcing the model to recompute later equipment, health, and battle outcomes. The benchmark includes 117 task categories and 159 question types, with the longest sequences containing about 1,200 events. It evaluates 14 model configurations, including GPT-5.5, Gemini 3.1 Pro, DeepSeek V4 Preview Pro, GLM-5.2, and Qwen3.5 Plus. In the hardest setting, the median score was only 11.3%, while the top-performing GPT-5.5-xhigh reached 59.3%. The team also said the dataset can be used for training: after adding 6,000 such samples, an internal A3B model improved on all eight external tests, gaining an average of 5.25 points.810
Anthropic2026-09-05 13:06:05Anthropic's Claude AI Converts Fermat's Last Theorem into 13 Million Lines of Self-Checking CodeAnthropic announced that its AI model Claude spent 11 days converting the 350-year-old mathematical problem Fermat's Last Theorem into a 13 million-line code proof. The proof can be self-verified by computers without relying on human trust. Anthropic stated that this achievement demonstrates AI's potential in formal verification and complex logical reasoning, and could potentially transform how mathematical proofs are conducted. The news was reported by Decrypt.760
OpenServ2026-07-24 05:35:17OpenServ and Neol Bring Enterprise AI Reasoning Into Real Production ConditionsOpenServ has entered a foundational design partnership with Neol to test and refine its AI reasoning framework in regulated, high-stakes production settings, with the resulting patterns now being folded into its platform.190